Tags: deepswe 1.1*

0 bookmark(s) - Sort by: Date ↓ / Title /

  1. Benjamin Marie writes that the effectiveness of an LLM in long-horizon agentic coding tasks depends heavily on the harness used to drive it rather than just the model itself. Through testing Qwen3.8 27B across three different interfaces—Mini-SWE Agent, Claude Code, and Pi—the author found that while specific configurations like "benchmaxxed" Pi can solve the highest number of tasks, other setups like Claude Code achieve better functional coverage (F2P). The study highlights how critical engineering choices, such as preserving reasoning traces or managing output token limits, are essential for successful agentic performance.

    - The evaluation used DeepSWE 1.1, a benchmark comprising 113 long-horizon tasks from 91 open-source repositories.
    - Performance varies significantly based on whether reasoning traces are preserved between turns and how context budgets are managed.
    - Pi at medium effort was found to offer the best balance of efficiency and accuracy.
    - Results were influenced by factors like session recovery, patch reliability, and output-token settings. author »

Top of the page

First / Previous / Next / Last / Page 1 of 0 SemanticScuttle - klotz.me: tagged with "deepswe 1.1"

About - Propulsed by SemanticScuttle